Papers with computational linguistics
Copied to clipboard
| Challenge: | COLING 2020 is the 28th International Conference on Computational Linguistics held online due to the COVID-19 pandemic. |
| Approach: | a volume of papers from the online system demonstration session was published . the volume was published by the international committee on computational linguistics . |
| Outcome: | The volume contains papers from the system demonstration session held online due to the COVID-19 pandemic. |
Copied to clipboard
| Challenge: | COLING 2018 is a three-hour tutorial series covering a range of core problems and exciting developments in computational linguistics and natural language processing. |
| Approach: | COLING 2018 has six tutorials covering a range of core problems and exciting developments in computational linguistics and natural language processing. |
| Outcome: | COLING 2018 will host six tutorials covering a range of core problems and exciting developments in computational linguistics and natural language processing. |
Copied to clipboard
| Challenge: | COLING 2018 is a conference on computational linguistics . the program committee accepted 35 papers out of 53 submissions based on quality of work and utility . |
| Approach: | a rigorous review process accepted 35 papers out of 53 submissions . the program committee consisted of 36 members and one chair from academia and industry . |
| Outcome: | the program committee accepted 35 papers out of 53 submissions . most accepted systems are user-interactive and feature rich graphical user interfaces based on the COLING 2018 program . |
Copied to clipboard
| Challenge: | COLING 2020 will be the first conference to have a dedicated track for computational linguistics research in real-world settings. |
| Approach: | COLING 2020 will feature a dedicated track for research related to computational linguistics deployed in real-world settings. |
| Outcome: | COLING 2020 will showcase commercially-driven research from diverse angles . an estimated 76% of submissions came from industry and 24% from academia . most submissions were from North America (44%), 27-28% from Europe and Asia, and 1% from Africa . |
Copied to clipboard
| Challenge: | Philip is a member of the Association for Computational Linguistics and is pursuing his PhD in computational linguistics. |
| Approach: | Philip is a member of the Association for Computational Linguistics and is pursuing a PhD in computational linguistics. |
| Outcome: | Philip is a member of the Association for Computational Linguistics and is pursuing two PhDs in computational linguistics and cognitive neuroscience. |
Copied to clipboard
| Challenge: | COLING 2022 theme is "NLP for the Grand Challenges of Our Time" COLing 2022 is the first conference of its kind in the world, and will be held in Gyeongju, Gyeonju . |
| Approach: | COLING 2022 theme is "NLP for the Grand Challenges of Our Time" a highly efficient committed team has worked for the organization of COLing 2022 . a hybrid COLTING is planned because of the unpredictable consequences of Covid-19 . |
| Outcome: | COLING 2022 theme is "NLP for the Grand Challenges of Our Time" a hybrid conference will be held because of the unpredictable consequences of Covid-19 . |
Copied to clipboard
| Challenge: | This tutorial provides background information for system developers and researchers working with Arabic in its various forms. |
| Approach: | This tutorial provides the necessary background information for working with Arabic in its various forms. |
| Outcome: | This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing. |
Copied to clipboard
| Challenge: | a tutorial on reproducibility in ML addresses the problem of research results that are not reproducible. |
| Approach: | They propose a tutorial to ensure reproducible research in ML with an emphasis on computational linguistics and NLP. |
| Outcome: | The proposed tutorial focuses on computational linguistics and NLP . it provides a framework for using reproducibility as a teaching tool in university-level computer science programs. |
Copied to clipboard
| Challenge: | Existing deep-learning approaches model code generation as text generation, but few of them account for compilability of the generated programs. |
| Approach: | They propose a three-stage pipeline utilizing compiler feedback for compilable code generation to improve compilability. |
| Outcome: | The proposed pipeline improves compilability of generated programs by combining compiler feedback, language model fine-tuning, and compilable discrimination. |
Copied to clipboard
| Challenge: | Currently, there is no widely accepted standard for evaluation of machine translation (MT) for Chinese-to-English translation, there are no standard for standardized training sets, development sets, and test sets. |
| Approach: | They propose to use Chinese-to-English machine translation as a benchmark . they build a highly competitive state-of-the-art MT system that outperforms reported results . |
| Outcome: | The proposed system outperforms reported results on NIST OpenMT test sets in almost all papers published in major conferences and journals in computational linguistics and artificial intelligence in the past 11 years. |
Copied to clipboard
| Challenge: | COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning. |
| Approach: | a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding . |
| Outcome: | This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 . |
Copied to clipboard
| Challenge: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
| Approach: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
| Outcome: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
Copied to clipboard
| Challenge: | This tutorial provides a historical overview of grounding and discusses its use in computational linguistics and in computational language processing. |
| Approach: | They introduce the concept of grounding and discuss future directions and open challenges . they will delve into recent progress in learning lexical semantics, syntax, and complex meanings through various forms of ground. |
| Outcome: | This course will provide an overview of the field of grounding and discuss future directions and challenges related to large language models and scaling. |
Copied to clipboard
| Challenge: | Using focus-background dichotomy, discourse and information structure of sentences are being studied in context. |
| Approach: | They propose to automate the analysis of focus in authentic written data by using a range of lexical, syntactic, and semantic features to achieve an accuracy of 78.1%. |
| Outcome: | The proposed approach achieves 78.1% accuracy for identifying focus in authentic written data. |
Copied to clipboard
| Challenge: | No existing methods can achieve effective text segmentation and word discovery in open domain Chinese texts. |
| Approach: | They propose a Bayesian-based method that can achieve effective text segmentation and word discovery in open domain. |
| Outcome: | The proposed method enjoys robust performance and transparent interpretation when no training corpus and domain vocabulary are available. |
Copied to clipboard
| Challenge: | Existing research indicates that disfluencies can constitute up to 5.9% of words in spontaneous speech, with repetitions accounting for over half of these disfluency. |
| Approach: | They propose to use a dataset to analyze reduplication and repetition in speech using computational linguistics to evaluate transformer-based models. |
| Outcome: | The proposed models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification. |
Copied to clipboard
| Challenge: | XPLAINSIM is a Python package that explains textual similarity in an easy-to-use way. |
| Approach: | They propose a Python package that unifies three approaches to explain text similarity . they demonstrate the value of the package through intuitive examples and empirical research . |
| Outcome: | XPLAINSIM is a Python package that unifies three approaches to explain text similarity . the authors show that the package is useful for explaining text similarities in a simple way . |
Copied to clipboard
| Challenge: | Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities. |
| Approach: | They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors. |
| Outcome: | The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors. |
Copied to clipboard
| Challenge: | Language of study extraction is an aspect of computational linguistics papers that is useful for analyses of trends and diversity in computational linguists. |
| Approach: | They propose to benchmark and evaluate automated language of study extraction from computational linguistics papers. |
| Outcome: | The proposed language extraction benchmarks show that they can extract languages from papers with accuracy without high computational costs. |
Copied to clipboard
| Challenge: | a novel approach for learning probabilistic context-free grammars from strings is proposed . strong learning means that there can be many structurally different PCFGs that define the same distribution over strings. |
| Approach: | They propose an algorithm that is a consistent estimator for a class of PCFGs that are anchored . they show that if the grammar is anchored, the parameters can be directly related to distributional properties of the anchoring strings. |
| Outcome: | The proposed algorithm is consistent for a large class of probabilistic context-free grammars . it shows that the proposed algorithm has good finite sample behavior . |
Copied to clipboard
| Challenge: | adverbs are the part of speech (POS) that has seen the least attention in computational linguistics due to its challenging nature. |
| Approach: | They propose to use Frame Semantics to characterize word meaning to uncover systematic gaps in adverb accounts. |
| Outcome: | The proposed approach can describe ambiguity, semantic roles, and null instantiation of adverbs. |
Copied to clipboard
| Challenge: | a new corpus-based study addresses racial stereotypes in social media conversations . a multilingual corpus of rhs is used to investigate how they are spread . |
| Approach: | They propose a corpus-based method for multilingual racial stereotype identification in social media conversational threads. |
| Outcome: | The proposed method sheds light on how racial hoaxes are spread and allows identification of negative stereotypes that reinforce them. |
Copied to clipboard
| Challenge: | Existing methods to disentangle an author's style from the content of their writing are limited by the reliance on human labels and the narrow focus of stylistic distinctions. |
| Approach: | They propose to use a surrogate task to learn authorship representations that are sensitive to writing style and to validate their hypothesis . |
| Outcome: | The proposed representations are sensitive to writing style and may be robust to topic drift over time. |
Copied to clipboard
| Challenge: | social media has brought with it a massive channel for spreading and reinforcing stereotypes . most stereotypes are expressed implicitly and identifying them automatically remains a challenge . |
| Approach: | They propose criteria to facilitate the subjective task of identifying the presence of stereotypes . they propose a newsCom-Implicitness corpus of 1,911 sentences, of which 426 are explicit and implicit racial stereotypes. |
| Outcome: | The proposed criteria show that they obtain different inter-annotator agreement values . the criteria are applied to a corpus of 1,911 sentences, of which 426 are explicit and implicit racial stereotypes . |
Copied to clipboard
| Challenge: | Existing studies on code-mixing have not been able to model human interactions in context. |
| Approach: | They propose to use a general-purpose code-mixing corpus to model human interactions and relationships in context while maintaining ethical standards. |
| Outcome: | The proposed corpus includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages. |
Copied to clipboard
| Challenge: | Romanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools. |
| Approach: | They introduce a Large Romanian Sentiment Data Set which is composed of 15,000 positive and negative reviews collected from the largest Romanian e-commerce platform. |
| Outcome: | The proposed data set is composed of 15,000 positive and negative reviews from the largest Romanian e-commerce platform. |
Copied to clipboard
| Challenge: | a new study of fear speech is under-resourced and fragmented. authors review existing definitions and propose a taxonomy that consolidates different dimensions of fear. |
| Approach: | They propose a taxonomy that consolidates different dimensions of fear for studying fear speech. |
| Outcome: | The proposed taxonomy consolidates different dimensions of fear for studying fear speech. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have enabled powerful new possibilities for semantic text analysis. |
| Approach: | They leverage GPT-4 to extract five semantic features from transcripts of spontaneous patient speech. |
| Outcome: | The proposed model significantly improves detection of AD in manually transcribed and automatically generated transcripts. |
Copied to clipboard
| Challenge: | Prior studies on identifying the existence or the type of complaints focus on building automatic classification models for identifying complaints. |
| Approach: | They propose to measure the intensity of complaints from text using Best-Worst Scaling method to estimate the popularity of posts on social media. |
| Outcome: | The proposed model can estimate the popularity of complaints on social media with best-worst scaling (BWS) method. |
Copied to clipboard
| Challenge: | Discussions on an appropriate annotation scheme for large and complex information are ongoing . multi-layer system allows a comprehensive description of relations between morphological properties, syntactic function and expressed meaning. |
| Approach: | They propose a multi-layer annotation scheme for the Prague Dependency Treebank . they propose morphological properties, syntactic function and expressed meaning as multi-layered systems . |
| Outcome: | The proposed scheme is sound and serves well for complex annotations. |
Copied to clipboard
| Challenge: | Using structured attention, a model can learn dialogue structure in unsupervised fashion. |
| Approach: | They propose to incorporate structured attention layers into a Variational Recurrent Neural Network model with discrete latent states to learn dialogue structure in an unsupervised fashion. |
| Outcome: | The proposed model learns semantic structures similar to templates used to generate a dialogue corpus on two-party datasets and on multi-party dialogues, disentangling dialogues without human annotation. |
Copied to clipboard
| Challenge: | a study of scientific topics and their evolution through time is proposed . we analyze scientific texts published in the field of computational linguistics . |
| Approach: | They propose a multidimensional approach to studying scientific topics through time and their relationships between them. |
| Outcome: | The proposed model analyzes scientific texts published in the ACL Anthology and compares them with case studies to understand how topics evolve and disappear over time. |
Copied to clipboard
| Challenge: | Structured representations have long been pivotal in computational linguistics, but their role remains ambiguous in the Large Language Models (LLMs) era. |
| Approach: | They propose a framework that integrates structured representations into LLMs from training-free and training-dependent perspectives. |
| Outcome: | The proposed framework integrates structured representations through natural language descriptions in LLM prompts while augmenting the model’s inference capability through fine-tuning on linguistically described structured representation. |
Copied to clipboard
| Challenge: | Current dialog systems require human experts to design the dialog structure, which is time consuming and sometimes insufficient to satisfy various customer needs. |
| Approach: | They propose to extract dialog structure using a modified VRNN model with discrete latent vectors. |
| Outcome: | The proposed model outperforms existing models on the ability to predict unseen data and is faster and more effective in a reinforcement learning setting. |
Copied to clipboard
| Challenge: | Complaining is a speech act used by humans to communicate a negative mismatch between reality and expectations . recent work on modeling complaints in natural language processing (NLP) has focused on distinguishing complaints from non-complaints in social media. |
| Approach: | They propose to classify complaints into various severity levels based on the face-threat that the complainer is willing to undertake and their purpose. |
| Outcome: | The proposed model achieves 55.7 macro F1 on binary complaint classification and 88.2 macro F1. |
Copied to clipboard
| Challenge: | a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text . |
| Approach: | They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags. |
| Outcome: | The proposed method detects humor in code-mixed tweets in English-Hindi. |
Copied to clipboard
| Challenge: | Humblebragging is a phenomenon in which individuals present self-promotional statements under the guise of modesty or complaints. |
| Approach: | They propose a task of automatically detecting humblebragging in text and propose '4-tuple definition' they also propose machine learning, deep learning, and large language models to perform the task . |
| Outcome: | The proposed model achieves an F1-score of 0.88 and is non-trivial even for humans. |
Copied to clipboard
| Challenge: | Existing methods for CB detection oversimplify the problem of CB as a binary classification task. |
| Approach: | They propose to use large language models to generate CB-related datasets . they propose to combine cognitive and linguistic models to help identify CB incidents . |
| Outcome: | The proposed approach aims to help researchers and policymakers make informed decisions . it uses large language models such as Claude-2 and Llama2-Chat to generate CB-related datasets . |
Copied to clipboard
| Challenge: | Existing computational models of the verbal morphology of the Métis language are insufficient to model the language's unique phonological interactions. |
| Approach: | They propose a finite-state computational model of the verbal morphology of Michif . they use composed finite state transducers to model concatenative morphologies . |
| Outcome: | The proposed model is based on a series of finite-state transducers. |
Copied to clipboard
| Challenge: | Existing studies on what constitutes a "basic" color term and its acquisition sequence are flawed . a pan-lingual approach may reveal general color trends more reliably than smaller datasets. |
| Approach: | They propose to operationalize and critique the Berlin and Kay color term hypotheses . they use 14 empirically-grounded computational linguistic metrics to analyze cross-linguistic data . |
| Outcome: | The proposed measures correlate strongly with the Berlin and Kay color term partition and their hypothesized universal acquisition sequence. |
Copied to clipboard
| Challenge: | Existing literature is agnostic about a parsing strategy of hierarchical models . a recent study showed that hierarchically model hierarchic structures capture grammatical dependencies much better than RNNs in targeted syntactic evaluations. |
| Approach: | They evaluated three LMs with head-final left-branching structures and Recurrent Neural Network Grammars with top-down and left-corner parsing strategies as hierarchical models. |
| Outcome: | The proposed model outperforms top-down and left-corner models against human reading times in Japanese. |
Copied to clipboard
| Challenge: | a large number of language models struggle to handle disfluencies, authors say . when a speaker hesitates, interrupts themselves, repeats or corrects words, or abandons phrases, it can make their speech fragmented. |
| Approach: | They propose to use disfluent queries to “clean” spontaneous speech . they propose to apply disfluencies to models that use different types of speech repairs . |
| Outcome: | The proposed model improves on a reading comprehension task using disfluent queries . the results suggest that disfluencies can improve model performance, rather than their removal . |
Copied to clipboard
| Challenge: | Bragging is a speech act employed to build a favorable self-image through positive statements about oneself. |
| Approach: | They propose to use tweets annotated for bragging to build a model that can predict bragging with macro F1 up to 72.42 and 35.95 for binary and multi-class bragging classification tasks respectively. |
| Outcome: | The proposed models predict bragging with macro F1 up to 72.42 and 35.95 in binary and multi-class classification tasks respectively. |
Copied to clipboard
| Challenge: | In this paper, we examine the problem of pejorative language, an under-explored topic in computational linguistics. |
| Approach: | They propose to automatically disambiguate pejorative usage in social media . they leverage online dictionaries to build a multilingual lexicon of pejorativ terms . |
| Outcome: | The proposed model can automatically disambiguate pejorative usage in social media posts . the proposed model is based on dictionaries and tweets . |
Copied to clipboard
| Challenge: | Figures of speech often deviate from their literal meanings to express deeper semantic implications. |
| Approach: | They propose a concept of figurative unit, which is the carrier of a figure, and build a Chinese corpus for Contextualized Figure Recognition. |
| Outcome: | The proposed model is based on 12 types of figures commonly used in Chinese . it shows that the proposed tasks are challenging for existing models . |
Copied to clipboard
| Challenge: | Dialog2Flow embeddings allow for modeling dialogs as continuous trajectories in a latent space with distinct action-related regions. |
| Approach: | They propose dialog2Flow embeddings that map dialogs to a latent space and cluster them according to their communicative and informative functions. |
| Outcome: | The proposed workflow embeddings show superior performance across domains. |
Copied to clipboard
| Challenge: | Online political advertising is an integral part of modern digital election campaigning. |
| Approach: | They propose to use textual and visual information from pre-trained neural models to infer the political ideology of an ad sponsor and identify whether the sponsor is an official political party or a third-party organization. |
| Outcome: | The proposed approach outperforms state-of-the-art methods for generic commercial ad classification and linguistic analysis to study the characteristics of political ads discourse. |
Copied to clipboard
| Challenge: | A human language is comprised of a pronunciation system and a writing system, both evolving and changing over time. |
| Approach: | They reformulate existing phonetic rules into a dataset of 70,943 entries for 17,001 Chinese characters and use it to perform a temporal prediction task. |
| Outcome: | The transformer-based model significantly advances the digitization and computational reconstruction of ancient Chinese phonology, providing a more complete and temporally contextualized resource for computational linguistics and historical research. |
Copied to clipboard
| Challenge: | a prior work using surprisal only considered within-sentence context, using n-grams, neural language models, or syntactic structure as conditioning context. |
| Approach: | They extend the surprisal approach to use broader topical context . they identify distinct patterns of neural activation for lexical surprised and topical surpresed . |
| Outcome: | The proposed method captures effects of local and topical contexts on processing . it shows that local and broad contextual cues recruit different brain regions . |
Copied to clipboard
| Challenge: | Recent work on the emergence of language between artificial agents has not isolated the effect of categorization power on inter-communication ability. |
| Approach: | They propose to use disentangled representations to quantify categorization power of agents to enable differential analysis between combinations of heterogeneous systems. |
| Outcome: | The proposed method reduces signaling accuracy by 40% despite encouraging compositionality in the artificial language. |
Copied to clipboard
| Challenge: | a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language. |
| Approach: | They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles . |
| Outcome: | The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy). |
Copied to clipboard
| Challenge: | a new paper aims to provide sentiment analysis tools for ancient languages . the current sentiment analysis resources only cover modern languages based on textual typologies . |
| Approach: | They propose to use manually-curated Latin lexicons to evaluate sentiment analysis tools . they propose a gold standard and a silver standard for evaluating lexical items . |
| Outcome: | The proposed lexicons are evaluated using a gold standard and a silver standard for sentiment analysis. |
Copied to clipboard
| Challenge: | a theoretical framework for low-resource parsing is understudied in computational linguistics but widely used in typological research . a novel approach uses Role and Reference Grammar to parse low-source languages . |
| Approach: | They propose to extend an existing RRG parser into a cross-lingual parsing model . they also adopt self-training to adapt the model to a related language with no trees . |
| Outcome: | The proposed model extends into a cross-lingual parser, and iteratively expands the training data. |
Copied to clipboard
| Challenge: | Informal social interaction is the primordial home of human language. |
| Approach: | They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future. |
| Outcome: | The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure. |
Copied to clipboard
| Challenge: | Pun generation aims to modify linguistic elements in text to produce humour or evoke double meanings. |
| Approach: | They propose to review pun generation datasets and methods across different stages . pun generation aims to produce humour or evoke double meanings . |
| Outcome: | This paper summarises both automated and human evaluation metrics used to assess the quality of pun generation. |
Copied to clipboard
| Challenge: | Using social networks, social media is a vital tool for emergency management and social media has been used to generate valuable information in crisis situations. |
| Approach: | They propose to measure for the first time the role of SA on urgency detection in tweets . they propose to use a two-layer annotation scheme to annotate tweets for both SA and urgency . |
| Outcome: | The proposed scheme combines two-layer annotation scheme and deep learning experiments to detect SA in a crisis corpus. |
Copied to clipboard
| Challenge: | Complaining is a basic speech act used to express a negative mismatch between reality and expectations in a particular situation. |
| Approach: | They present a systematic analysis of complaints in computational linguistics . they collect annotated data set of written complaints expressed on Twitter . |
| Outcome: | The proposed model achieves predictive performance of up to 79 F1 using distant supervision. |
Copied to clipboard
| Challenge: | Existing Natural Language Inference (NLI) datasets are not related to scientific text. |
| Approach: | They propose a large dataset for NLI that captures the formality in scientific text and contains 107,412 sentence pairs extracted from scholarly papers on NLP and computational linguistics. |
| Outcome: | The proposed model achieves a Macro F1 score of only 78.18% and an accuracy of 78.23%. |
Copied to clipboard
| Challenge: | Personality profiling has long been used in psychology to predict life outcomes. |
| Approach: | They present the trajectory of automatic personality detection from purely psychology approaches to the latest purely natural language processing approaches on large social media datasets. |
| Outcome: | The proposed models have been compared with the most recent approaches on large social media datasets. |
Copied to clipboard
| Challenge: | Structured data, such as database tables or XML trees, often contain short natural language labels that describe the data structure itself or provide content (attribute values). Conventional NLP tools, such supervised sequence labellers or embeddings trained on full sentences, do not perform well on structured data. |
| Approach: | They propose to design a type of abbreviated grammar that is called the Language of Data and to investigate the grammatical properties of such labels. |
| Outcome: | The proposed model outperforms models trained on standard text on tokenisation, part-of-speech tagging, and named entity recognition over real-world structured data. |
Copied to clipboard
| Challenge: | a recent increase in quantitative studies of scientific text collections has led to a significant increase in the use of semantic labeling techniques. |
| Approach: | They propose to use semantic class labels to enhance a well-known resource . they use semantic labels to assign semantic class labeling to technical terms . |
| Outcome: | The proposed approach enhances the ACL Anthology Reference Corpus with semantic class labels for 20,000 technical terms . the goal is to use this information as one feature in the profiling of scientific papers, communities, and disciplines. |
Copied to clipboard
| Challenge: | WikiDragon is a Java Framework designed to give developers in computational linguistics an intuitive API to build, parse and analyze instances of MediaWikis. |
| Approach: | They introduce WikiDragon, a Java Framework that allows developers to build, parse and analyze instances of MediaWikis on their computers. |
| Outcome: | The framework is based on the Wikipedia, Wiktionary, WikiSource or WikiNews and evaluates link extraction, diachronic network analysis and the impact of different frameworks to text analysis. |
Copied to clipboard
| Challenge: | ACL OCL is a scholarly corpus derived from the ACL Anthology . it provides metadata, PDF files, citation graphs and additional structured full texts . |
| Approach: | They present ACL OCL, a scholarly corpus derived from the ACL Anthology . it integrates metadata, PDF files, citation graphs and additional structured full texts . they highlight how it applies to observe trends in computational linguistics . |
| Outcome: | The ACL OCL spans seven decades and contains 73,285 papers . the scholarly corpus is based on the ACL Anthology and is available from HuggingFace . |
Copied to clipboard
| Challenge: | Replicability and reproducibility are core ideas of modern scientific methods. |
| Approach: | They describe challenges encountered in reproducing the results of a top performing system in computational linguistics. |
| Outcome: | The proposed system was able to reproduce the results of a task 7 in the domain of natural language processing and computational linguistics. |
Copied to clipboard
| Challenge: | Polysemy is the phenomenon where a single word form possesses two or more related senses. |
| Approach: | They propose an unsupervised framework to quantify polysemy for words in multiple languages . they use syntactic knowledge to infuse the framework with syntaktic knowledge . |
| Outcome: | The proposed framework is based on syntactic knowledge and is compared with existing methods in English, French and Spanish. |
Copied to clipboard
| Challenge: | Semantic Textual Similarity (STS) is a key indicator of the encoding capabilities of embedding models. |
| Approach: | They propose to use Pearson’s correlation coefficient as a loss function to refine model performance beyond contrastive learning to achieve a Spearman’s ceiling. |
| Outcome: | The proposed method surpasses state-of-the-art strategies with minimal amount of fine-grained annotated samples. |
Copied to clipboard
| Challenge: | Currently, there is no publicly available corpus for diachronic text analysis due to the lack of accurate temporal metadata. |
| Approach: | They propose to add missing temporal metadata to the Gutenberg corpus by using open web, Wikipedia, and Open Library API sources. |
| Outcome: | The proposed corpus includes 53,774 books with a total of 3.8 billion tokens in 11 languages, produced between 1600 and 2000. |
Copied to clipboard
| Challenge: | Existing AV techniques, including stylometric and deep learning, face limitations in terms of data requirements and lack of explainability. |
| Approach: | They propose a technique that leverages Large-Language Models (LLMs) to provide step-by-step stylometric explanation prompts to verify authorship. |
| Outcome: | The proposed technique outperforms state-of-the-art baselines, operates effectively with limited training data, and enhances interpretability through intuitive explanations. |
Copied to clipboard
| Challenge: | Existing methods for embedding words from colexification networks are limited to the word level, ignoring lexical relations that would only hold for parts of words in a given language. |
| Approach: | They propose to embed concepts from automatically constructed colexification networks . they use lexical similarity ratings and word association data to evaluate the methods . |
| Outcome: | The proposed methods capture and represent different semantic relationships between concepts. |
Copied to clipboard
| Challenge: | Existing work describes paragraph-level counter-argument generation task as paragraph-based . however, sentence-level generation can be quite different due to its unique constraints and brevity-focused challenges. |
| Approach: | They propose a benchmark framework for sentence-level counter-argument generation . they use an annotated debate forum dataset to generate high-quality counter-argments . |
| Outcome: | The proposed framework and evaluator are competitive in counter-argument generation tasks. |
Copied to clipboard
| Challenge: | Existing constituency treebanks are limited in out-of-domain settings, therefore constituency parsing is still a challenge. |
| Approach: | They propose a novel method for constituency parsing using large language models . they use a cross-domain constituency treebank to fill missing words with the incomplete one . |
| Outcome: | The proposed method achieves state-of-the-art performance on average compared with baselines on five target domains of MCTB. |
Copied to clipboard
| Challenge: | The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which is used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. |
| Approach: | They propose to use a dataset of sentences manually annotated for sentiment to train a robust sentiment identifier for parliamentary proceedings. |
| Outcome: | The proposed model performs very well on languages not seen during fine-tuning and additional fine- tuning data from other languages significantly improves the target parliament’s results. |
Copied to clipboard
| Challenge: | Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research. |
| Approach: | They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse. |
| Outcome: | The proposed dataset is the largest of its kind that is publicly accessible. |
Copied to clipboard
| Challenge: | Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed. |
| Approach: | They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries. |
| Outcome: | The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards. |
Copied to clipboard
| Challenge: | Training language models and examining their linguistic behaviors is a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing. |
| Approach: | They replicate three prior studies with hyperparameters varied within a practical range and show that modest hyperparametric changes can alter qualitative conclusions about models’ linguistic abilities. |
| Outcome: | The results show that hyperparameter changes can alter qualitative conclusions and reverse the ranking of models. |
Copied to clipboard
| Challenge: | Among the approximately 7,000 languages spoken globally, fewer than 20 receive substantial attention in NLP research. |
| Approach: | They propose to use African multi-modal speech and text data to validate African multimodal models and validate them on targeted language data. |
| Outcome: | The African Languages Lab's results show that the proposed model outperforms untrained models in 31 languages and a 1B-parameter model beats the commercial system in Yoruba and Twi. |